Eval-to-Evolution with skill-upper
Create evals through natural conversation, then let skill-upper diagnose failures, repair or expand cases, and rerun skill-up until the suite evolves.
Learn more
Measure Skill quality, then turn failures into automatic eval fixes and the next iteration.

skill-up combines evaluation and evolution for Agent Skills. It makes quality measurable through declarative YAML evals, isolated multi-engine runs, flexible judges, and structured reports. Then skill-upper turns failures into progress: it can automatically repair or expand the eval suite and rerun the loop with you.
The same workflow runs locally or in CI, supports Anthropic evals.json imports, and emits JSON, JUnit, and HTML reports.
